This work propose an audio feature pipeline to support machine learning tasks through the extraction of the Mel Frequency Cepstral Coefficients and Mel-spectrogram which is then used as the input of an Convolutional Neural Network which is trained to make the classification tasks. This approach enables de creation of a rich-feature dataset and an end-to-end pipeline that reduces the gap between the audio and Machine Learning ready models with application in sound classification, speech recognition and spatial audio analysis.
Loading....